Goto

Collaborating Authors

 safety standard


OpenAI Cancels Its Advanced Model Launch, Then Heads to Lunch With Trump

Mother Jones

Russian trolls push fake videos with our name on them in order to sow chaos ahead of the election. Trump hates scrutiny, so he cries "FAKE NEWS" and bans the free press from his briefings. And AI companies steal our reporting in the pursuit of power and profit. We're ready to keep uncovering the truth that those in power don't want you to know. But we need your support to do that.


OpenAI Delays Release of Latest Model Over Safety Concerns

WIRED

The company said its latest Astra model would undergo more work to meet safety standards, and issued an apology for the way it handled the hacking of an Australian government website. OpenAI has cancelled plans to release its latest GPT-6.1 Astra system next month after the model failed to meet safety standards. Research and safety leaders decided not to ship the model after finding it was worse at sticking to human users' values and goals than previous systems, OpenAI told WIRED. "It didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done," head of safety systems Saachi Jain said. The company said it has other new models coming soon which do meet its safety standards and plans to release other Astra models in future. The agent accessed non-public data, ran commands, and wrote files onto the server.


How Self-Driving Cars Might Change Crash Testing Forever

WIRED

Lie down in a car, and you're more likely to get seriously hurt in a crash. But a driverless future is forcing governments to revisit long-standing safety issues. In a video posted on X, a Tesla enthusiast hops into what was once the driver's seat of a new matte gold Cybercab . "Once" because the vehicle, which officially hit the streets earlier this month as part of Tesla's Robotaxi ride-hail service, doesn't have a steering wheel or brake pedals . For the person in the front seat, this opens up new possibilities.


Anthropic's CEO proposes a three-step plan to curb AI development

Engadget

At least one CEO of an AI company is calling for a slower pace when it comes to developing artificial intelligence. Anthropic's CEO Dario Amodei wrote a lengthy post detailing a goal of pacing the speed at which AI is built, instead of forging ahead at the rapid rate that AI is currently on. Amodei proposed a three-tiered approach to achieve this, with Anthropic already committing to the first step. The first measure calls for "frontier AI" companies to commit to "ongoing, employee-like access" for third-party evaluators who would focus on verifying compliance with certain safety standards, evaluating if AI model training is aligned with the goal of slowing down and reporting incidents. The second step requires these AI companies to establish "common safety standards" with the help of governments in order to limit the rate of unchecked AI progress.


Tesla's Cybercab Officially Launches Today. It's Already Under Investigation

WIRED

The US government is investigating whether the Cybercab, which lacks a steering wheel and pedals, meets vehicle safety standards. Tesla's Cybercab, a distinctive two-seater without a steering wheel or brake pedals, is set to start picking up members of the public in Austin, Texas, this evening. But the vehicle is already under investigation by the US federal government, which is probing whether it meets federal safety standards. The investigation comes just hours after the electric automaker welcomed hundreds of fans to downtown Austin to ride in the driverless Cybercabs. Tesla plans to deploy the vehicles on its Robotaxi ride-hail network, which is currently operating in a handful of cities in Texas and Florida.


An Ebike Company Was Sued for Misleading Info on Safety. It Points to a Big Problem

WIRED

An Ebike Company Was Sued Over Alleged Safety Misinformation. Several Chinese businesses associated with the ebike brand Aipas settled a lawsuit with Amazon and the safety certification company UL in July. Amazon and UL, a company that tests and certifies that products meet safety standards, settled a lawsuit in July with several Chinese companies that manufacture ebikes under the Aipas brand name. The US companies accused the Chinese firms in January of claiming in Amazon listings and on Aipas' website that the ebikes had been certified by UL, giving recognizable, brand-name assurance to online shoppers worried about risks like ebike lithium-ion battery fires . On July 15, a federal judge signed off on a permanent injunction in which the companies agreed that they would no longer use the UL mark.


I Met With China's Top AI Experts. They're Freaking Out, Too

WIRED

The AI arms race between China and the US has researchers on both sides worried about a "Chernobyl moment." Just over a week ago, I attended a major artificial intelligence conference in Zhongguancun, Beijing's bustling high-tech district. It was packed with fascinating sessions touching on everything from recursive self-improvement--the idea that models can tweak their own code and advance indefinitely--to humanoid robots. And it featured a few legends of computing, including Whitfield Diffie, co-inventor of public-key cryptography, and Andrew Barto, who won the Turing Award with Rich Sutton for his pioneering work on reinforcement learning. But I left with one takeaway above all else: The US and China should put their fierce AI rivalry to the side.


Canada moves to ban social media for children under 16 and regulate AI chatbots

The Japan Times

Several countries have been considering tightening rules around AI use as well as social media use for children. OTTAWA - The Canadian government introduced a digital safety bill on Wednesday that would ban social media for children under 16 with exemptions for platforms that meet certain safety standards, months after Australia enacted the world's first social media ban for young people. The bill also aims to make AI chatbots safer by setting up a digital regulator to establish safety standards, a government official said. Companies could face penalties of 3% of global revenue or up to 10 million Canadian dollars ($7.2 million), whichever is more, for failing to comply. "Social media platforms and AI chatbots are designed to capture attention. They do not support healthy childhood development and have become a source of anxiety, isolation, depression and a range of other mental health challenges for many young Canadians," said Marc Miller, minister of Canadian identity and culture.


Illinois Lawmakers Just Passed America's Strongest AI Safety Bill

WIRED

Illinois Lawmakers Just Passed America's Strongest AI Safety Bill The bill requires companies like OpenAI, Anthropic, and Google to have third parties confirm they're following safety standards. The Illinois House of Representatives passed a bill on Wednesday requiring frontier AI labs like OpenAI, Anthropic, and Google DeepMind to have their safety practices audited by a third party. If signed into law, AI safety experts tell WIRED, it would be the nation's leading check on the power of major AI companies . The bill, SB 315, now heads to governor JB Pritzker's desk. In a post on social media on Wednesday, Pritzker said he plans to sign the bill, citing a need to hold Big Tech accountable.


Protect: Towards Robust Guardrailing Stack for Trustworthy Enterprise LLM Systems

arXiv.org Artificial Intelligence

The increasing deployment of Large Language Models (LLMs) across enterprise and mission-critical domains has underscored the urgent need for robust guardrailing systems that ensure safety, reliability, and compliance. Existing solutions often struggle with real-time oversight, multi-modal data handling, and explainability -- limitations that hinder their adoption in regulated environments. Existing guardrails largely operate in isolation, focused on text alone making them inadequate for multi-modal, production-scale environments. We introduce Protect, natively multi-modal guardrailing model designed to operate seamlessly across text, image, and audio inputs, designed for enterprise-grade deployment. Protect integrates fine-tuned, category-specific adapters trained via Low-Rank Adaptation (LoRA) on an extensive, multi-modal dataset covering four safety dimensions: toxicity, sexism, data privacy, and prompt injection. Our teacher-assisted annotation pipeline leverages reasoning and explanation traces to generate high-fidelity, context-aware labels across modalities. Experimental results demonstrate state-of-the-art performance across all safety dimensions, surpassing existing open and proprietary models such as WildGuard, LlamaGuard-4, and GPT-4.1. Protect establishes a strong foundation for trustworthy, auditable, and production-ready safety systems capable of operating across text, image, and audio modalities.